SoK:AI 增强型二进制逆向工程知识系统化研究
文章背景与核心概要
二进制逆向工程是软件理解、漏洞发现、恶意软件调查和固件审计的核心环节,但由于编译过程中语义信息的丢失,它始终具有极大的挑战性。近年来,机器学习、大语言模型(LLMs)以及智能体(Agentic)AI 系统的飞速发展,推动了 AI 增强型二进制逆向技术的大规模应用。然而,现有的相关研究在逆向领域、代码表示方法、学习策略以及评估标准上高度分散。
本文针对 2015 年以来发表的 246 篇研究论文,开展了首个关于 AI 增强型二进制逆向工程的全面知识系统化(SoK)研究。作者将这些研究系统梳理为 22 个不同的二进制逆向推理领域,并引入了一个连接传统与 AI 增强型逆向流水线的统一分类法。该框架将传统的分析技术、二进制衍生人工制品、表示策略、学习范式与下游推理任务紧密串联,同时阐明了 LLMs 和智能体 AI 的演进角色。最终,本研究揭示了底层通用结构、弥补了评估空白,并为构建基于证据且具备实际部署能力的下一代 AI 逆向系统指明了未来的研究方向。
SoK: AI-Augmented Binary Reversing
SoK: AI-Augmented Binary Reversing
Summary
Summary
Binary reversing is a critical process for software understanding, vulnerability discovery, malware investigation, and firmware auditing, though it remains inherently challenging due to semantic information loss during compilation. While recent advancements in machine learning, large language models (LLMs), and agentic AI systems have driven the adoption of AI-augmented binary reversing, the existing body of work is heavily fragmented across reversing domains, representations, learning approaches, and evaluation methods.
This paper provides the first comprehensive Systematization of Knowledge (SoK) on AI-augmented binary reversing by analyzing 246 research papers published since 2015. The authors organize these works into 22 distinct binary reversing inference domains and introduce a unified taxonomy that bridges conventional and AI-augmented reversing pipelines. This structured framework connects traditional analysis techniques, binary-derived artifacts, representation strategies, learning paradigms, and downstream inference tasks while highlighting the evolving roles of LLMs and agentic AI. Ultimately, the study uncovers common underlying structures, addresses evaluation gaps, and outlines future research pathways for building evidence-grounded and deployable AI-augmented reversing systems.
Document Metadata
Document Metadata
Field Details arXiv ID arXiv:2606.17398 [cs.CR] Subjects Cryptography and Security ( cs.CR); Artificial Intelligence (cs.AI); Software Engineering (cs.SE)ACM Classes D.4.6; K.6.5 Submission History • v1: June 16, 2026
• v2: September 4, 2026 (This version)Document Stats 21 pages, 7 tables, 4 figures
Authors
Authors
- Yujeong Kwon
- Yiyue Zhang
- Kexin Pei
- Dokyung Song
- Hyungjoon Koo
Abstract
Abstract
Binary reversing is fundamental to software understanding, vulnerability discovery, malware investigation, and firmware auditing. However, it remains inherently challenging due to the lossy transformation of semantic information during compilation. Recent advances in machine learning, large language models (LLMs), and agentic AI systems have accelerated the adoption of AI-augmented binary reversing. Yet, the resulting body of work has become increasingly fragmented across reversing domains, artifact representations, learning approaches, and evaluation practices.
This paper presents the first comprehensive systematization of knowledge on AI-augmented binary reversing. We collect 246 research papers published since 2015, and organize them into 22 binary reversing domains according to the inference tasks. We further introduce a unified taxonomy spanning conventional and AI-augmented reversing pipelines. Our taxonomy connects traditional analysis techniques, binary-derived artifacts, representation strategies, learning paradigms, and downstream inference tasks, while clarifying the emerging roles of LLMs and agentic AI systems.
By establishing a common vocabulary and structured framework, we offer a holistic view of the field's evolution over the past decade. Our study reveals common structures underlying seemingly disparate approaches, highlights persistent technical challenges and evaluation gaps, and identifies promising opportunities for future research. Collectively, these insights clarify the current state of the field and provide a foundation for the next generation of evidence-grounded and practically deployable AI-augmented binary reversing systems.
Resources & Links
Resources & Links
- Full-Text Options: View PDF | HTML Version | TeX Source
- License: Creative Commons Attribution 4.0 International
- External Citations: Google Scholar | Semantic Scholar | NASA ADS
